Papers with sentence encoders
Extracting Text Representations for Terms and Phrases in Technical Domains (2023.acl-industry)
Copied to clipboard
| Challenge: | Large pre-trained language models are extensively used in modern NLP systems. |
| Approach: | They propose an unsupervised approach to encoding using character-based models and pre-trained sentence encoders to reconstruct large pre-trained embedding matrices. |
| Outcome: | The proposed approach matches the quality of sentence encoders in technical domains and is 5 times smaller and up to 10 times faster on high-end GPUs. |
Evaluating Composition Models for Verb Phrase Elliptical Sentence Embeddings (N19-1)
Copied to clipboard
| Challenge: | ellipsis is a natural language phenomenon where part of a sentence is missing and its information must be recovered from its context. |
| Approach: | They develop models for embedding VP-elliptical sentences using word embeddments . they extend existing verb disambiguation and sentence similarity datasets to elliptic phrases . |
| Outcome: | The proposed models outperform existing models on verb disambiguation and sentence similarity datasets and their linear counterparts. |
Testing Paraphrase Models on Recognising Sentence Pairs at Different Degrees of Semantic Overlap (2023.starsem-1)
Copied to clipboard
| Challenge: | Existing models for paraphrase detection are not suitable for many applications . existing datasets ignore and fail to test models in this setup . |
| Approach: | They propose to use adversarial paradigms to test paraphrase detection models . they propose to examine the sensitivity to different degrees of semantic overlap . |
| Outcome: | Empirical results show that paraphrase models and different sentence encoders appear successful on evaluations, but measuring the degree of semantic overlap remains a big challenge for them. |
Multi-Source Text Classification for Multilingual Sentence Encoder with Machine Translation (2024.naacl-srw)
Copied to clipboard
| Challenge: | Pre-trained multilingual sentence encoders suffer from performance degradation for non-English languages. |
| Approach: | They propose a method of machine translating a source sentence into English and then inputting it together with the source sentence in a multi-source manner. |
| Outcome: | The proposed method improves the performance of pre-trained multilingual sentence encoders in Japanese on sentiment analysis and topic classification tasks. |
SmartMatch: Real-Time Semantic Retrieval for Translation Memory Systems (2026.eacl-demo)
Copied to clipboard
Ernesto L. Estevanell-Valladares, Salima Lamsiyah, Alicia Picazo-Izquierdo, Tharindu Ranasinghe, Ruslan Mitkov, Rafael Muñoz
| Challenge: | Translation Memory (TM) systems are core components of computer-aided translation tools . however, they fail to retrieve semantically relevant content when surface similarity is low. |
| Approach: | They propose an open-source demo and evaluation toolkit for TM retrieval that connects modern sentence encoders and strong lexical/fuzzy baselines with a vector database. |
| Outcome: | The proposed toolkit exposes the end-to-end retrieval pipeline through a web-based UI for qualitative inspection and preference logging. |
Discrete Cosine Transform as Universal Sentence Encoder (2021.acl-short)
Copied to clipboard
| Challenge: | Modern sentence encoders capture underlying linguistic characteristics of words . Discrete Cosine Transform (DCT) is an efficient alternative to averaging . |
| Approach: | They propose to use a Discrete Cosine Transform to generate universal sentence representations in different languages. |
| Outcome: | The proposed model captures the underlying syntactic characteristics of a given text without compromising practical efficiency. |
Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations (D19-1)
Copied to clipboard
| Challenge: | Prior work on pretrained sentence embeddings and benchmarks focused on the capabilities of stand-alone sentences. |
| Approach: | They propose a test suite of tasks to evaluate whether sentence representations include broader context information. |
| Outcome: | The proposed training objectives help to encode different aspects of information in document structures. |
On Measuring Social Biases in Sentence Encoders (N19-1)
Copied to clipboard
| Challenge: | Word embeddings such as word2vec and GloVe exhibit human-like implicit biases based on gender, race, and other social constructs. |
| Approach: | They propose a simple generaliza test to measure bias in word embeddings by comparing two sets of target-concept words to two sets . |
| Outcome: | The proposed test shows that word2vec and word2Ve exhibit human-like implicit biases based on gender, race, and other social constructs. |
Exploiting Semantics in Neural Machine Translation with Graph Convolutional Networks (N18-2)
Copied to clipboard
| Challenge: | Semantic representations have long been argued as potentially useful for enforcing meaning preservation and improving generalization performance of machine translation methods. |
| Approach: | They propose to integrate semantic representations into neural machine translation by injecting a semantic bias into sentence encoders and achieving improvements in BLEU scores. |
| Outcome: | The proposed representations achieve better BLEU scores over the linguistic-agnostic and syntax-aware versions on the English–German language pair. |
Robust Fragment-Based Framework for Cross-lingual Sentence Retrieval (2021.findings-emnlp)
Copied to clipboard
Nattapol Trijakwanich, Peerat Limkonchotiwat, Raheem Sarwar, Wannaphong Phatthiyaphaibun, Ekapol Chuangsuwanich, Sarana Nutanong
| Challenge: | Cross-lingual Sentence Retrieval (CLSR) aims at retrieving parallel sentence pairs that are translations of each other from a multilingual set of comparable documents. |
| Approach: | They propose a framework for cross-lingual sentence retrieval that uses a collection of fragments to improve sentence retrievals. |
| Outcome: | The proposed framework improves the retrieval robustness of the base sentences encoded by m-USE, LASER, and LaBSE. |
Towards Improving Adversarial Training of NLP Models (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Recent methods for generating NLP adversarial examples involve combinatorial search and expensive sentence encoders for constraining the generated instances. |
| Approach: | They propose to use vanilla adversarial training to train NLP models using a word substitution attack optimized for vanilla adversary training. |
| Outcome: | The proposed approach improves model performance and standard accuracy and can defend against other types of word substitution attacks. |
ConvFiT: Conversational Fine-Tuning of Pretrained Language Models (2021.emnlp-main)
Copied to clipboard
Ivan Vulić, Pei-Hao Su, Samuel Coope, Daniela Gerz, Paweł Budzianowski, Iñigo Casanueva, Nikola Mrkšić, Tsung-Hsien Wen
| Challenge: | Existing Transformer-based language models (LMs) are not effective as sentence encoders when used off-the-shelf. |
| Approach: | They propose a method which turns a pretrained LM into a universal conversational encoder and task-specialised sentence encoder. |
| Outcome: | The proposed framework achieves state-of-the-art ID performance across the board with particular gains in the most challenging, few-shot setups. |
Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic Representations (2024.naacl-long)
Copied to clipboard
Sihao Chen, Hongming Zhang, Tong Chen, Ben Zhou, Wenhao Yu, Dian Yu, Baolin Peng, Hongwei Wang, Dan Roth, Dong Yu
| Challenge: | Sentence embeddings are typically learned to recognize the semantic relation between two text inputs. |
| Approach: | They introduce a contrastively-learned contextual embedding model for fine-grained semantic representation of text. |
| Outcome: | The proposed model is able to produce contextual embeddings corresponding to different atomic propositions, i.e. semantic equivalence between propositions across different text sequences. |
Multilingual Representation Distillation with Contrastive Learning (2023.eacl-main)
Copied to clipboard
| Challenge: | Contextual representations from large pretrained language models encode semantic information from two or more languages. |
| Approach: | They integrate contrastive learning into multilingual representation distillation and use it for quality estimation of parallel sentences. |
| Outcome: | The proposed model outperforms existing models with similarity searches and filtering tasks across low-resource languages. |
PTEB: Towards Robust Text Embedding Evaluation via Stochastic Paraphrasing at Evaluation Time with LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing evaluations of sentence embedding models rely on static tests like the Massive Text Embedding Benchmark (MTEB) repeated tuning on a fixed suite can inflate reported performance and obscure real-world robustness. |
| Approach: | They propose a dynamic protocol that generates meaning-preserving paraphrases at evaluation time and aggregates results across multiple runs. |
| Outcome: | The proposed protocol generates meaning-preserving paraphrases at evaluation time and aggregates results across multiple runs. |
Bitext Mining Using Distilled Sentence Representations for Low-Resource Languages (2022.findings-emnlp)
Copied to clipboard
| Challenge: | a new study aims to extend multilingual representation learning beyond the hundred most frequent languages . current work on multilingual sentence representations has focused on training one model which handles all languages of interest . |
| Approach: | They propose a teacher-student approach to extend existing monolingual sentence embedding space to new languages. |
| Outcome: | The proposed model outperforms the original LASER encoder in 44 African languages . the model can be used to train multiple languages and learn new languages if they have the same training data . |
DialogueCSE: Dialogue-based Contrastive Learning of Sentence Embeddings (2021.emnlp-main)
Copied to clipboard
| Challenge: | Conventional approaches to learning sentence embeddings from dialogues employ the siamese-network for this task, but such architecture yields a large gap between training and evaluating. |
| Approach: | They propose a dialogue-based contrastive learning approach to learn sentence embeddings from dialogues using a siamese-network. |
| Outcome: | The proposed model outperforms baseline methods on three multi-turn dialogue datasets in terms of MAP and Spearman’s correlation measures. |
ConveRT: Efficient and Accurate Conversational Representations from Transformers (2020.findings-emnlp)
Copied to clipboard
| Challenge: | ConveRT is a pretraining framework for conversational AI that is computationally heavy, slow, and expensive to train. |
| Approach: | They propose a pretraining framework for conversational tasks that is efficient, lightweight, and inexpensive. |
| Outcome: | The proposed model achieves state-of-the-art performance across widely established responses . it trains substantially faster than existing state- of-the art models . |
SentEval: An Evaluation Toolkit for Universal Sentence Representations (L18-1)
Copied to clipboard
| Challenge: | a toolkit for evaluating the quality of universal sentence representations is available for download and preprocessing . word embeddings are not trained to perform well on one specific task, but their value lies in their transferability . evaluation of general-purpose word and sentence embeddables has been problematic . |
| Approach: | They propose a toolkit to evaluate the quality of universal sentence representations. |
| Outcome: | The proposed toolkit includes scripts to download and preprocess datasets and an easy interface to evaluate sentence encoders. |
Neural Extractive Summarization with Hierarchical Attentive Heterogeneous Graph Network (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing extractive summarization methods focus on balancing salience and redundancy between sentences. |
| Approach: | They propose a hierarchical attentive heterogeneous graph for text summarization that models sentences . they propose to iteratively refine the sentence representations and deliver the labels by message passing . |
| Outcome: | The proposed method outperforms existing extractive summarization methods on large corpus. |
Continual Learning for Sentence Representations Using Conceptors (N19-1)
Copied to clipboard
| Challenge: | Existing sentence encoders for distributed representations of sentences are limited in their performance on fixed corpora. |
| Approach: | They propose a continual learning scenario for distributed representations of sentences . they initialize sentence encoders with corpus-independent features and update them sequentially . |
| Outcome: | The proposed sentence encoder can learn features from new corpora while maintaining its competence on previously encountered corporales. |
Word Reordering for Zero-shot Cross-lingual Structured Prediction (2021.emnlp-main)
Copied to clipboard
| Challenge: | Current sentence encoders are word order sensitive, resulting in poor performance . Adapting word order from one language to another is key in cross-lingual structured prediction. |
| Approach: | They propose a new module to organize words following the source language order . they build structured prediction models with bag-of-words inputs and introduce a module to do this . |
| Outcome: | The proposed model significantly improves target language performance for languages that are distant from the source language. |
Towards Structure-aware Paraphrase Identification with Phrase Alignment Using Sentence Encoders (2022.coling-1)
Copied to clipboard
| Challenge: | Existing paraphrase identification datasets exhibit high correlation between positive pairs and the degree of their lexical overlap. |
| Approach: | They propose to combine sentence encoders with an alignment component by representing each sentence as a list of predicate-argument spans and decomposing the sentence-level meaning comparison into the alignment between their spans. |
| Outcome: | The proposed approach improves performance and interpretability for various sentence encoders. |
Intermediate Self-supervised Learning for Machine Translation Quality Estimation (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for machine translation quality estimation (QE) rely on annotated data. |
| Approach: | They propose a self-supervised learning task for machine translation (MT) that orients a pre-trained model towards the target task. |
| Outcome: | The proposed method outperforms existing methods on English-to-German and English- to-Russian translation directions and is comparable to existing models. |
ALIGN-SIM: A Task-Free Test Bed for Evaluating and Interpreting Sentence Embeddings through Semantic Similarity Alignment (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Sentence embeddings play a pivotal role in a wide range of NLP tasks . evaluating and interpreting these dense vectors remains an open challenge to date . |
| Approach: | They propose a task-free test bed for evaluating and interpreting sentence embeddings . they examined five classical and eight LLM-induced sentence embedders based on semantic similarity alignment criteria . |
| Outcome: | The proposed test bed consists of five semantic similarity alignment criteria . it shows that none of the embeddings aligned with the criteria compared to other benchmarks . |
Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling (P19-1)
Copied to clipboard
Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman
| Challenge: | State-of-the-art models in natural language processing (NLP) often incorporate sentence encoder functions which generate a sequence of vectors intended to represent the in-context meaning of each word in an input text. |
| Approach: | They conduct the first large-scale systematic study of candidate pretraining tasks, comparing 19 different tasks as alternatives and complements to language modeling. |
| Outcome: | The proposed model can be used to train sentences on language modeling tasks. |
Sentence Meta-Embeddings for Unsupervised Semantic Textual Similarity (2020.acl-main)
Copied to clipboard
| Challenge: | Existing word embeddings combine complementary strengths of their components to achieve unsupervised semantic similarity (STS). |
| Approach: | They propose to ensemble pre-trained sentence encoders into sentence meta-embeddings to achieve unsupervised Semantic Textual Similarity (STS) they adapt dimensionality reduction, generalized Canonical Correlation Analysis and cross-view auto-encoders to their work. |
| Outcome: | The proposed method achieves 3.7% to 6.4% Pearson’s r over single-source word embeddings on the STS Benchmark and on the StS12-STS16 datasets. |
Map of Encoders – Mapping Sentence Encoders using Quantum Relative Entropy (2026.acl-long)
Copied to clipboard
| Challenge: | a method to compare and visualise sentence encoders at scale is proposed . we map encoder LLMs using QRE-based feature vectors, which are then projected to 2D . |
| Approach: | They propose a method to compare and visualise sentence encoders at scale by creating a map of encoder . they construct a QRE-based map of sentences covering 1101 publicly available sentence encoded sentences . |
| Outcome: | The proposed method compares sentence encoders at scale by creating a map of encoder models . it shows that the map accurately reflects relationships between encoder and unit base encoder . |
Contrastive Learning-based Sentence Encoders Implicitly Weight Informative Words (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Embedding a sentence into a point in a highdimensional continuous space plays a foundational role in the natural language processing. |
| Approach: | They propose to use contrastive loss to fine-tune sentences by inverse word frequency . they also show that more informative words receive greater weight than less informative ones . |
| Outcome: | The proposed method improves the performance of sentence embeddings by weighing them based on information-theoretic quantities. |
A Simple Geometric Method for Cross-Lingual Linguistic Transformations with Pre-trained Autoencoders (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have used probing tasks to verify the presence of linguistic properties in vector representations, but it is unclear whether they can be manipulated to indirectly steer them. |
| Approach: | They validate a geometric mapping technique to transform linguistic properties without tuning . they use a pre-trained multilingual autoencoder to transform three linguistic property . |
| Outcome: | The proposed method can be used without tuning of the pre-trained autoencoder . the results are validated in monolingual and cross-lingual settings . |
Ranking-Enhanced Unsupervised Sentence Representation Learning (2023.acl-long)
Copied to clipboard
Yeon Seonwoo, Guoyin Wang, Changmin Seo, Sajal Choudhary, Jiwei Li, Xiang Li, Puyang Xu, Sunghyun Park, Alice Oh
| Challenge: | Unsupervised sentence representation learning has progressed through contrastive learning and data augmentation methods such as dropout masking. |
| Approach: | They propose a novel unsupervised sentence encoder, RankEncoder, which predicts the semantic vector of an input sentence by leveraging its relationship with other sentences in an external corpus. |
| Outcome: | The proposed unsupervised sentence encoder achieves 80.07% Spearman’s correlation, a 1.1% improvement over the previous state-of-the-art system. |
Varying Sentence Representations via Condition-Specified Routers (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing sentences cannot account for different aspects of semantic similarity between two sentences. |
| Approach: | They propose a transformer-style framework that generates conditioned sentences . they propose 'conditional' STS, which measures similarity between two sentences based on condition sentences - a task that requires a sentence embedding model capable of generating distinct representations for the same sentence under different conditions. |
| Outcome: | The proposed framework is superior to existing models on two condition sentences . it can generate conditioned sentences while maintaining model parameters and computational efficiency . |
Reimagining Intent Prediction: Insights from Graph-Based Dialogue Modeling and Sentence Encoders (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to intent prediction are limited in highly specialized fields, such as closed-domain dialogue systems, where context comprehension is of paramount importance. |
| Approach: | They propose a method that uses scenario dialog graphs to model dialogues as sequences of transitions between intents, representing distinct goals or requests. |
| Outcome: | The proposed method significantly advances the field of dialogue systems, providing valuable insights into the effectiveness and potential limitations of the proposed approaches. |